Papers with Multilingual large language models
Thank You, Stingray: Multilingual Large Language Models Can Not (Yet) Disambiguate Cross-Lingual Word Senses (2025.findings-naacl)
Copied to clipboard
Samuel Cahyawijaya, Ruochen Zhang, Jan Christian Blaise Cruz, Holy Lovenia, Elisa Gilbert, Hiroki Nomoto, Alham Fikri Aji
| Challenge: | Existing studies on multilingual large language models have raised concerns about their reliability beyond English. |
| Approach: | They propose a benchmark for cross-lingual sense disambiguation that uses false friends to identify the limitation of cross-linguistic sense disembarrassment in LLMs. |
| Outcome: | The proposed benchmark pinpoints the limitation of cross-lingual sense disambiguation in LLMs by using false friends in four languages. |
Concept Space Alignment in Multilingual LLMs (2024.emnlp-main)
Copied to clipboard
| Challenge: | Multilingual large language models generalize somewhat across languages, but it is unclear whether this is a result of improved, implicit alignment, or of something else, e.g., linguistic overlap or semi-parallel subsets of training data. |
| Approach: | They hypothesize that implicit alignment is the reason for generalization in multilingual large language models. |
| Outcome: | The proposed model generalizes well across languages, but lacks linearity. |
AlignX: Advancing Multilingual Large Language Models with Multilingual Representation Alignment (2025.emnlp-main)
Copied to clipboard
| Challenge: | Multilingual large language models (LLMs) possess impressive multilingual understanding and generation capabilities, but performance and cross-lingual alignment often lag for non-dominant languages. |
| Approach: | They propose a representation-level framework to enhance multilingual performance of pre-trained LLMs by integrating multilingual semantic alignment and language feature integration. |
| Outcome: | The proposed framework improves multilingual capability of pre-trained LLMs by bringing representations closer and improving cross-lingual alignment. |
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models? (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing practices of fine-tuning and evaluating multilingual large language models may not align with this objective due to a heavy reliance on translation. |
| Approach: | They propose to use translated or native instruction data to fine-tune multilingual large language models. |
| Outcome: | The proposed model can be fine tuned and evaluated in multilingual large language models . the results show that native or translated data can be used to compare model performance . |
Pruning Multilingual Large Language Models for Multilingual Inference (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Multilingual large language models (MLLMs) demonstrate better zeroshot learning performance in non-English languages compared to large language model trained on English-dominant data. |
| Approach: | They propose a pruning approach to prune large language models using bilingual sentence pairs from English and other languages to enhance their performance in non-English language. |
| Outcome: | The proposed pruning strategy enhances the MLLMs’ performance in non-English language. |
Paths Not Taken: Understanding and Mending the Multilingual Factual Recall Pipeline (2025.emnlp-main)
Copied to clipboard
| Challenge: | Multilingual large language models (LLMs) exhibit factual inconsistencies across languages . authors identify two primary sources of error: insufficient engagement of reliable English-centric mechanism for factual recall, and incorrect translation from English back into the target language for the final answer. |
| Approach: | They propose two vector interventions to redirect the model toward better internal paths for higher factual consistency. |
| Outcome: | The proposed interventions increase the recall accuracy by over 35 percent for the lowest-performing language. |
MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation datasets lack cross-lingual alignment, leaving assessments of multilingual capabilities fragmented in both language and skill coverage. |
| Approach: | They propose to use multilingual consistency as a complementary metric to assess performance bottlenecks and guide model improvement. |
| Outcome: | The proposed model lacks cross-lingual alignment and language coverage gaps between state-of-the-art models. |
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches address key factors that influence multilingual ICL, but they do not integrate them into the model. |
| Approach: | They propose a method that quantifies and optimally balances three factors for improved example selection. |
| Outcome: | Experiments on mCSQA and TYDI show that the proposed method outperforms existing methods. |
Explainability and Interpretability of Multilingual Large Language Models: A Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing literature on multilingual large language models lacks transparency in their internal processes. |
| Approach: | They propose to use multilingual large language models to examine their explainability and interpretability methods. |
| Outcome: | The present study examines the explainability and interpretability of multilingual large language models. |
Error Analysis of Multilingual Language Models in Machine Translation: A Case Study of English-Amharic Translation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Multilingual large language models have significantly advanced machine translation, yet challenges remain for low-resource languages like Amharic. |
| Approach: | They evaluated the performance of NLLB-200 and M2M in English-Amharic bidirectional translation using the Lesan AI dataset. |
| Outcome: | The proposed models outperformed the existing models in English-Amharic bidirectional translation using the Lesan AI dataset. |
Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs (2026.acl-long)
Copied to clipboard
Guy Mor-Lan, Omer Goldman, Matan Eyal, Adi Mayrav Gilady, Sivan Eiger, Idan Szpektor, Avinatan Hassidim, Yossi Matias, Reut Tsarfaty
| Challenge: | Multilingual large language models have minimized the fluency gap between languages, but they are exposed to the risk of biases as knowledge and norms may propagate across languages. |
| Approach: | They propose a test set with 2,156 questions in 12 languages to quantify models' biases . they show a global bias towards answers relevant to the US-locale . |
| Outcome: | The proposed model can answer locale-ambiguous questions in 12 languages. |
The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining (2026.acl-long)
Copied to clipboard
Jiandong Shao, Raphael Tang, Crystina Zhang, Karin Sevegnani, Pontus Stenetorp, Jianfei Yang, Yao Lu
| Challenge: | Existing research suggests that multilingual large language models can achieve impressive cross-lingual understanding despite largely monolingual pretraining. |
| Approach: | They compare a monolingual-only corpus with a standard web corpus that removes all multilingual documents and then retrain the models from scratch under controlled conditions. |
| Outcome: | The results show that removing bilingual data causes translation performance to drop 56% in BLEU, whereas code-switching contributes minimally. |
Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian Languages (2026.acl-long)
Copied to clipboard
| Challenge: | Multilingual large language models are expensive to pretrain and suffer from imbalances across languages and datasets. |
| Approach: | They propose a family of Indian language-only autoregressive language models trained on open-source language-specific data for the five most spoken Indian languages. |
| Outcome: | The proposed model outperforms most larger models up to 8B across all five languages. |
From Representation to Choice: Tracing Decision Emergence Across Languages in LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in large language models have made them highly multilingual, but how they internally reason remains unexplored. |
| Approach: | They propose to model multilingual reasoning through a decision-making perspective using aligned multiple-choice questions from the mMMLU benchmark. |
| Outcome: | The proposed model shows that languages share similar activation spaces, but subtle divergences emerge as decisions propagate through transformer layers. |